Papers with video-language multimodal models

1 papers
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing video-to-text summarization evaluation methods depend heavily on human-written reference summaries.
Approach: They propose a reference-free metric evaluating candidate summaries directly against source videos through multimodal question answering.
Outcome: The proposed metric assesses candidate summaries directly against source videos through multimodal question answering.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations